Back

Aperture Neuro

Organization for Human Brain Mapping

Preprints posted in the last 30 days, ranked by how well they match Aperture Neuro's content profile, based on 20 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.

1
Reinforcement Learning via Brain Feedback for real-time fMRI-based adaptive stimulus generation

Gallitto, G.; Englert, R.; Kincses, B.; Kotikalapudi, R.; Li, J.; Hoffschlag, K.; Ali, S.; Bingel, U.; Spisak, T.

2026-08-13 bioinformatics 10.64898/2026.08.08.743648 medRxiv
Top 0.1%
6.1%
Show abstract

Traditional fMRI studies rely on predefined task paradigms, where fixed stimulus designs limit the flexibility with which brain-stimulus relationships can be explored. Here, we introduce Reinforcement Learning via Brain Feedback (RLBF), a framework and open-source software package for adaptive stimulus optimization using real-time fMRI. RLBF reverses the conventional direction of inference by using neural responses to guide the exploration of stimulus spaces through reinforcement learning, enabling optimization of predefined brain targets such as regional activity or multivariate neural signatures. The accompanying Python-based software provides a modular framework integrating real-time fMRI data processing, reinforcement learning agents, adaptive stimulus generation, simulation-based testing, and experiment monitoring. Its flexible architecture allows researchers to customize preprocessing pipelines, reward functions, stimulus spaces, and RL strategies for diverse closed-loop neuroimaging applications. We validate the framework in a proof-of-concept study (N=10), demonstrating real-time optimization of a simple visual stimulus space by adapting checkerboard contrast and frequency to maximize primary visual cortex (V1) responses within a single 10-minute fMRI session. RLBF provides an extensible foundation for brain-guided stimulus optimization and enables new approaches for investigating neural specificity, individualized brain-stimulus relationships, and adaptive experimental design.

2
Network- and Measure-Specific Mid-Term Reliability of Multi-Echo Resting-State Functional Magnetic Resonance Imaging on a Compact 3 Tesla Scanner

Kang, D.; Welker, K. M.; Hermes, D.; Bernstein, M. A.; Huston, J.; Shu, Y.

2026-08-13 neuroscience 10.64898/2026.08.07.743542 medRxiv
Top 0.1%
5.4%
Show abstract

1.IntroductionUnderstanding mid-term test-retest reliability and within-subject variability is important for interpreting changes observed in longitudinal and intervention studies. The reliability of resting-state functional magnetic resonance imaging (rs-fMRI) is known to vary across measures and brain regions. However, how reliability differs across functional networks and connectivity-and amplitude-based measures, and whether multi-echo acquisition and processing modify these patterns, remain incompletely characterized. MethodsTwenty-two healthy volunteers underwent two rs-fMRI sessions 15.7 {+/-} 4.0 days apart on a Compact 3T scanner. Multi-echo, middle-echo, and independently acquired single-echo datasets were compared, with multi-echo independent component analysis additionally evaluated as a denoising approach. Functional connectivity (FC) and three amplitude-based measures were evaluated using the Schaefer 400 parcellation. Reliability was systematically assessed using intraclass correlation coefficient (ICC), within-subject standard deviation (wSD), and systematic bias at edge or regional, and network levels. ResultsAcquisition-dependent differences in reliability were generally modest. Multi-echo acquisition and processing increased functional connectivity strength and the magnitude of amplitude-based measures and improved inferior cortical coverage, but these enhancements did not consistently translate into substantially higher ICC or lower wSD. In contrast, reliability showed clear network-dependent differences. FC reliability varied markedly across network pairs and was not explained by connectivity strength alone; pairs involving the default mode and control networks generally showed more favorable profiles than several somatomotor and visual network pairs. Fractional amplitude of low-frequency fluctuations (fALFF) also showed network-dependent reliability, with the most favorable regional reproducibility observed in the default mode and control networks and lower reproducibility in the somatomotor and visual networks. ConclusionThese findings provide practical mid-term reliability benchmarks for rs-fMRI on a Compact 3T scanner and show that measurement stability varies more clearly across measures and functional networks than across acquisition approaches. Key pointsO_LIMid-term test-retest reliability varied more clearly across resting-state measures and functional networks than across acquisition and processing approaches. C_LIO_LIMulti-echo acquisition and processing enhanced functional connectivity strength, amplitude-based signal magnitude, and inferior cortical coverage but did not consistently improve reliability. C_LIO_LIFunctional connectivity strength and fractional amplitude of low-frequency fluctuations showed distinct network-specific reliability profiles, with more favorable reproducibility in default mode and control networks than in several somatomotor and visual networks. C_LI

3
Altered Speech Processing in Childhood Listening Difficulties as Revealed by Chirped Speech Event-Related Potentials

Petley, L.; Wicks, T.; Miller, L. M.; Blankenship, C.; Chatwin, J.; Bormann, B. M.; Whittle, R. S.; Moore, D. R.

2026-08-17 otolaryngology 10.64898/2026.08.13.26360392 medRxiv
Top 0.1%
4.1%
Show abstract

Objective: Impaired understanding of noisy or degraded speech is a central feature of listening difficulties (LiD), but the possible causes of these symptoms are wide-ranging. Accordingly, recent research underscores the need to study these deficits using a test battery approach. Event-related potentials are useful objective metrics for studying LiD, but probing function across the speech processing hierarchy using traditional protocols is sequential and unrealistic in clinical settings. The novel chirped speech (Cheech) method combines natural speech with acoustic chirps to overcome these limitations. This study examines its utility for profiling childhood LiD. Methods: Twenty-eight children (15 typically developing, 13 with LiD), aged 8-17 years old, listened to a 17-minute Cheech story and detected a target word within the story via button press while EEG data were collected from 53 scalp sites. Results: Cheech successfully evoked responses from the auditory brainstem response through to the brain's language centers, as reflected by the N400 effect. Unlike TD children, those with LiD demonstrated N400 effects with atypical distributions that favored frontal rather than the typical parietal sites. A trend towards a delayed and reduced amplitude Wave V was also observed. Conclusions: Hierarchical examination of speech processing using Cheech primarily implicates altered language processing as a contributing factor to LiD, with the frontal topography of the N400 effect for those with LiD potentially suggesting a greater reliance on deliberate memory retrieval during the speech perception task. Significance: LiD could arise due to auditory and/or cognitive factors. The present results demonstrate the feasibility of objective, parallel measurement across this hierarchy and point to impaired language processing as a possible mechanism.

4
ALFIE: Anatomy-aware enhancement of Low FIEld 64mT T2-weighted neonatal brain MRI for structural analysis

Cawley, P.; Uus, A.; Colford, K.; Padormo, F.; Teixeira, R.; Tomazinho, I.; UNITY Consortium, ; Williams, S. C. R.; Edwards, A. D.; O'Muircheartaigh, J.; Arichi, T.; Hajnal, J. V.; Rutherford, M. A.

2026-08-28 pediatrics 10.64898/2026.08.25.26361317 medRxiv
Top 0.1%
4.0%
Show abstract

Purpose: To develop and evaluate an anatomy-aware deep learning framework for enhancement of neonatal 64mT T2-weighted MRI that improves anatomical visibility while preserving native ultra-low-field contrast and enabling quantitative structural analysis. Methods: A multitask network, jointly performing image enhancement and tissue segmentation, was trained on 75 and evaluated on 20 paired neonatal 64mT/3T MRI datasets spanning a broad range of gestational ages and pathologies. To preserve native 64mT contrast, 3T images were locally harmonized before training. The framework also generated quality-control maps and regional volumetric measurements. Volumetric agreement was further assessed in 40 paired term-born control datasets. Results: Enhanced 64mT images showed improved image quality metrics and better delineation of cortical, deep gray matter, ventricular, white matter, and posterior fossa structures while maintaining native contrast characteristics. Tissue segmentations demonstrated good agreement with reference 3T labels. Volumetric measurements showed excellent correspondence with 3T across major tissue compartments, with only small systematic regional biases. Conclusions: Anatomy-aware enhancement enables automated tissue segmentation and volumetric analysis directly from neonatal 64mT MRI while preserving native image contrast. These findings support the feasibility of quantitative neonatal neuroimaging at ultra-low field.

5
10.5 Tesla High-Resolution Macaque Brain MRI for In vivo and Ex vivo Connectivity Studies

Warrington, S.; Selim, M. K.; Tendler, B. C.; Moeller, S.; Farooq, H.; Wu, W.; Pisharady, P. K.; Adriany, G.; Auerbach, E. J.; Folloni, D.; Bratch, A.; Manea, A. M.; Grafft, T.; Jungst, S.; Harel, N.; Waks, M.; Pestilli, F.; Yacoub, E.; Lenglet, C.; Ugurbil, K.; Heilbronner, S. R.; Miller, K. L.; Jbabdi, S.; Zimmermann, J.; Sotiropoulos, S. N.

2026-08-26 neuroscience 10.64898/2025.12.22.695917 medRxiv
Top 0.1%
3.3%
Show abstract

Mapping brain connectivity in primates remains a major challenge due to difficulties in resolving microscopic white matter architecture, while maintaining whole-brain coverage. Increasing imaging spatial resolution is key for disambiguating fibre configurations within smaller anatomical volumes. Here, we present novel developments that allow high-resolution diffusion MRI of the macaque brain using one of the world's highest-field human MRI scanners operating at 10.5 Tesla, allowing both in vivo and ex vivo macaque brain imaging. Our approach achieves very high spatial resolution across both tissue states, (up to 580 m)3 in vivo and (300 m)3 ex vivo, with diffusion weighting up to b = 6000 s/mm2. We detail methodological advances in data acquisition, image reconstruction, processing and whole-brain tractography that overcome critical challenges associated with ultra-high-field imaging. This work establishes a new framework for high-resolution in vivo and ex vivo neuroimaging of the NHP brain at 10.5 T using a human bore scanner, paving the way for subsequent analyses of brain connectivity across species and tissue states at unprecedented detail. The dataset, along with all processing pipelines, containerised workflows, and reusable web services, is openly shared to support reproducibility and future integration with microscopy for studying white matter microstructure and connections at the mesoscale.

6
AFIDs-Validator: An Open-Access AI-Guided Platform for Learning Anatomical Landmark Placement

Taha, A.; Bansal, D.; Kai, J.; Kuehn, T.; Stanley, O. W.; Park, P.; Thurairajah, A.; Snyder, M.; Gilmore, G.; Abbass, M.; Mahmoudian, B.; Liu, V. M.; Thrower, J.; Khan, A. R.; Lau, J. C.

2026-08-24 scientific communication and education 10.64898/2026.08.20.746086 medRxiv
Top 0.1%
3.3%
Show abstract

Accurate localization of anatomical landmarks is a foundational skill in anatomy and imaging that is often taught informally through expert mentorship, requiring access to data and desktop software. There is no openly accessible, interactive resource that teaches neuroanatomy with quantitative feedback. We present the AFIDs-Validator (validator.afids.io), an open-access, browser-based platform that pairs guided instruction with quantitative assessment. The platform combines (1) a learning mode in which a language-model neuroanatomy tutor operates inside an MRI viewer, giving anatomy-first instruction that responds to the learner's current image slice, orientation, and cursor position; and (2) a validation engine that accepts a learner's landmark file and returns per-landmark Euclidean error against expert-annotated references spanning 21 brain templates. To make the feedback interpretable, we analyzed 15,000 landmark annotations across 132 human subjects and found that landmark difficulty varies fourfold (median error ranged from 0.37 mm at the anterior commissure to 1.50 mm at the temporal horns) with heavy-tailed distributions at every landmark. These distributions are compiled into per-landmark reliability priors, so learners are scored against the empirical spread of trained raters rather than an arbitrary threshold, and difficult landmarks are not mistaken for poor performance. The AFIDs-Validator requires no installation, licensed software, or local data, and all code, reference data, and tutor design are openly released.

7
A framework for quality assurance in human intracranial electrophysiology

Herz, N.; Cao, R.; Qiu, S.

2026-08-12 neuroscience 10.64898/2026.08.06.743130 medRxiv
Top 0.1%
3.2%
Show abstract

Intracranial electroencephalography (iEEG) provides an unprecedented opportunity to directly record neural activity and causally perturb the human brain through electrical stimulation. Yet, the increasingly collaborative nature and complexity of modern iEEG studies pose substantial challenges for experimental control, data quality, and standardization. Unlike most experimental modalities, human iEEG data are acquired within dynamic clinical environments, where patient condition, recording quality, hardware configuration, and experimental protocols may vary across recording sessions and collaborating sites. The resulting heterogeneity creates opportunities for technical and procedural failures that often remain undetected until downstream analyses, when corrective action is no longer possible. Here, we present a framework for standardized session-level quality assurance in human iEEG research and provide an open-source implementation compatible with Brain Imaging Data Structure (BIDS)-organized datasets. The framework defines four complementary domains of quality assessment crucial for human iEEG studies: protocol fidelity, behavioral integrity, stimulation validation, and signal quality. These domains integrate electrophysiological recordings, behavioral event logs, and stimulation metadata to verify data completeness, confirm participant engagement, validate stimulation delivery, and identify potentially compromised recording channels. Automated quality metrics and standardized diagnostic visualizations are generated following each testing session, enabling rapid identification of technical and procedural failures while corrective action is still possible. By providing a standardized approach to session-level quality assurance, the framework improves data integrity, enhances reproducibility, facilitates analyst training, and supports harmonized data collection across laboratories and clinical sites.

8
A multi-b-value test-retest diffusion MRI brain dataset for model validation and reproducibility assessment

Pieciak, T.; Guadilla, I.; Ciupek, D.; Navarro-Gonzalez, R.; Merino-Caviedes, S.; Villacorta-Aylagas, P.; Magdaleno Humayor, L.; Villa Aparicio, M.; Rueda-Ramos, J.; Santiesteban Mendo, R.; Moro Boyero, R.; Tristan Vega, A.

2026-08-27 neuroscience 10.64898/2026.08.23.746449 medRxiv
Top 0.1%
3.2%
Show abstract

Transparent assessment of diffusion magnetic resonance imaging (dMRI) techniques with empirical verification of confounding factors requires adequately designed protocols and collected datasets. Publicly available diffusion-weighted MR datasets often provide limited sampling across b-values, making it difficult to study optimal acquisition protocols or the relationships between different processes occurring in brain tissue. In this work, we introduce a new densely sampled longitudinal test-retest diffusion-weighted MR dataset of the brain. Our dataset was collected from eleven healthy volunteers, each scanned four times: two sessions on consecutive days, which form the test data, followed by two additional sessions completed one week later (retest data). The data were acquired using twenty-two b-values ranging from 10 to 3000 s/mm2, along with structural T1-weighted scans. Potential applications of the dataset include, but are not limited to, assessing longitudinal reproducibility and reliability of quantitative metrics, evaluating robust and outlier-resistant estimation techniques, investigating experimental factors affecting estimation procedures, and verifying optimal acquisition protocols for different signal models. The dataset is publicly available in raw and fully preprocessed variants.

9
From Channel-Pair Connectivity to Brain Networks: An Open Graph Theoretical Pipeline for fNIRS Hyperscanning

Moshe, Y. H.; Sharma, M.; Dahan, A.; Gvirts, H.

2026-08-28 neuroscience 10.64898/2026.08.25.746918 medRxiv
Top 0.1%
3.2%
Show abstract

Despite the growing use of functional near-infrared spectroscopy (fNIRS) hyperscanning to record brain activity simultaneously from interacting individuals in naturalistic settings, most analyses quantify functional connectivity separately for each channel pair. The resulting collection of pairwise estimates is difficult to integrate into a network-level characterization of intra- and inter-brain organization. Here, we present an open, configuration-driven Python toolkit that transforms preprocessed fNIRS hyperscanning time series into functional connectivity graphs. The toolkit constructs a bipartite inter-brain network for each dyad and separate intra-brain networks for each participant, computes node- and graph-level measures, and exports adjacency matrices, edge lists, analysis-ready summary tables, reproducibility metadata, and standardized visualizations. Dataset-specific parameters, including directory structure, participant naming, channel selection, epoch extraction, and edge-retention criteria, are defined in a human-readable YAML configuration file, enabling the same workflow to accommodate differently organized datasets without changes to the source code. We illustrate the pipeline using a representative recording from a mother-infant fNIRS hyperscanning dataset and present the resulting network outputs. The toolkit provides a reproducible framework for moving from pairwise functional connectivity estimates to network-level analyses of dyadic and individual brain organization.

10
BRAIN CAST: An MRIQC-guided pipeline for age- and sex-specific pediatric brain MRI template construction, validated by downstream structural fidelity

Hu, Y.; Contreras-Vidal, J. L.

2026-08-07 neuroscience 10.64898/2026.08.02.742256 medRxiv
Top 0.1%
2.7%
Show abstract

Pediatric neuroimaging needs age- and sex-appropriate references, yet existing atlases span broad age ranges that blur development or lack sex specificity. We present BRAIN CAST: 28 year-by-year, sex-specific brain MRI templates covering ages 5-18, built from 1,272 quality-screened children in the Healthy Brain Network by an MRIQC-guided pipeline combining reduced-strength denoising, cerebrospinal-fluid-anchored intensity normalization, deep-learning skull stripping and iterative groupwise diffeomorphic registration. We evaluate templates not by image sharpness, which is not comparable across intensity conventions, but by the structural bias they induce downstream. Held-out children align to their matched template with sub-voxel gray-white interface error (1.1 mm); on a direction-symmetric surface-distance metric BRAIN CAST matches the best single-template reference and outperforms an age-specific pediatric atlas in 189 of 189 subjects. Female cortex is fit measurably better by female than by male templates, an effect no sex-neutral reference can provide. Templates, tissue-probability maps and the containerized pipeline are released.

11
Biomarker Fidelity Score - A Quantitative Framework for Individual-Level Validation of Explainability Methods in 3D Alzheimer's Disease MRI Classification

Lepcha, D. C.; Ali, A.; Martin, S. A.; Syed-Abdul, S.

2026-08-20 neuroscience 10.64898/2026.08.15.744687 medRxiv
Top 0.1%
2.5%
Show abstract

Explainability methods applied to deep learning models for Alzheimer's disease neuroimaging produce attribution maps that vary substantially across methods and architectures, yet no validated quantitative framework exists for determining which method most faithfully localises attribution signal within established AD biomarker anatomy at the individual subject level. Existing validation approaches rely on group-level comparisons or qualitative visual inspection, leaving individual-level biomarker alignment uncharacterised. We introduce the Biomarker Fidelity Score (BFS), a quantitative tool measuring spatial overlap between individual-level 3D explainability attention maps and atlas-registered AD-relevant neuroimaging ROIs across thirteen anatomically defined structures including hippocampus, entorhinal cortex, amygdala, and parahippocampal gyrus. Five explainability methods (GradCAM++, Integrated Gradients, DeepSHAP, LRP, ScoreCAM) were benchmarked across three volumetric architectures (3D ResNet-18, DenseNet-121, Swin-UNETR) on 327 balanced ADNI-3 subjects. Integrated Gradients achieved the highest BFS across all architectures while GradCAM++ consistently showed the lowest biomarker alignment (all p<0.001, Friedman test). The complete BFS pipeline replicated these rankings without retraining on 207 independent OASIS-3 subjects, with maximum absolute difference of 0.0005 across all fifteen method-architecture combinations and Spearman rank correlation of 0.964 between cohort rankings. By offering an externally validated, individual-level, biomarker-grounded quantitative standard, BFS equips clinicians and AI developers with practical guidance for selecting trustworthy explainability methods in AD neuroimaging.

12
Testing the reliability of novel Voxel Placement approaches for Magnetic Resonance Spectroscopy

Chhabra, H.; Hehl, M.; Cuypers, K.; Dydak, U.; Nitsche, M. A.; Genc, E.; Burke, M.

2026-08-21 neuroscience 10.64898/2026.08.11.744164 medRxiv
Top 0.2%
2.2%
Show abstract

BackgroundSingle-voxel magnetic resonance spectroscopy (MRS) is a non-invasive method for measuring clinically and cognitively relevant metabolites. Reliable measurements require precise voxel placement across sessions and participants. We developed a scanner-console-based approach to improve voxel placement precision. MethodsIn a crossover design (n=7; six sessions each), we compared test-retest reliability of three voxel placement methods in a reference benchmark (left parietal cortex) and a technically challenging region (left ventromedial prefrontal cortex). Methods included (1) conventional anatomy-based placement, (2) mask-guided real-time positioning (MGRP), and (3) semiautomated session-locked voxel repositioning (SSVR). Resting-state MRS data were acquired using PRESS and MEGA-PRESS. Within-subject reliability of voxel placement and metabolite concentrations, namely, total N-acetylaspartate (tNAA), total Creatine (tCr), GABA (gamma-aminobutyric acid), and Glx (glutamate + glutamine) are reported using the coefficient of variation (CV), the intraclass correlation coefficient (ICC), minimal detectable change (MDC), and the spatial overlap. ResultsSSVR markedly improved voxel placement reliability, increasing spatial overlap (up to 88%) and achieving near-perfect geometric reproducibility (ICC = 0.99) compared to conventional anatomy-based placement and MGRP. SSVR improved tissue composition consistency and reduced metabolite variability in the technically challenging region (variability reduction of [~]70% tCr, [~]59% tNAA, and [~]51% Glx) while further refining already stable measurements in the benchmark region (tNAA from [~]15% to [~]10%). ConclusionBoth MGRP and SSVR improved voxel placement and metabolite measurement reproducibility compared with conventional anatomy-based placement. SSVR further enhanced within-subject reproducibility across repeated sessions, particularly in the technically challenging region, providing a robust approach for longitudinal single-voxel MRS studies.

13
Developing an open-source framework for LLM evaluation of patients using EHR clinical documentation; performance of LLMs relative to medical professionals

Barrett, L.; Joshi, N.; North, A. S.; Dimitrov, L.; Maughan, E. F.; Ross, T.; Pankhania, R.; Paramjothy, K.; Minty, I.; Farache-Trajano, L.; Smith, S. L.; Mason, K. A.; Bhargava, E. K.; Donnelly, C.; Fatoum, H.; Padiyar, A.; Kader, Z.; Chan, C. H. K.; Schilder, A. G.; Mehta, N.

2026-08-24 otolaryngology 10.64898/2026.08.21.26361031 medRxiv
Top 0.2%
2.1%
Show abstract

Background: Large language models (LLMs) have shown increasing capability in medical knowledge tasks, yet how they perform in extracting structured clinical information from real-world clinical documentation remains uncertain. We evaluated the performance of LLMs relative to medical professionals in extracting SNOMED-coded clinical information from openly available Ear, Nose and Throat (ENT) EHRs from MTSamples, examining both reliability and accuracy metrics. Methods: We evaluated the performance of seven LLMs (including GPT-4o, Claude 3.5, Gemini 1.5 Pro, Gemma 3 and three LLAMA variants) against annotations from fourteen medical professionals who served as both study authors and data annotators. Each annotator independently extracted seven categories of clinical information from 98 publicly available ENT clinical documents: socio-demographics, symptoms, signs, diagnoses, treatments, risk factors, and test results. Standardised medical terminology was enforced through SNOMED-CT code assignment, enabling standardised comparison through Cohen's Kappa. We employed Bayesian hierarchical modelling to test non-inferiority of medic-LLM agreement compared to medic-medic agreement, using Beta distributed likelihood functions with weakly informative priors. Non-inferiority margins of 0.05, 0.10, and 0.15 were assessed with 95% posterior probability thresholds. Results: Cohen's Kappa for inter-rater reliability was 0.752 (95% CI: 0.710 - 0.794) between medical professionals and 0.391 (95% CI: 0.362-0.420) between LLMs and medical professionals. Bayesian analysis showed medic-medic agreement (posterior mean 0.813, 95% CI: 0.755-0.860) exceeded medic-LLM agreement (0.659, 95% CI: 0.633-0.684) by 0.154 (95% CI: 0.091-0.209). Non-inferiority was rejected at all tested margins (delta = 0.05, 0.10, 0.15). Agreement varied by clinical category, with smallest differences for test results and largest for diagnoses. GPT-4o achieved 97.0% precision and 84.9% recall, with a 7.5% false positive rate. Conclusions: Current LLMs do not achieve inter-rater reliability levels comparable to medical professionals in clinical information extraction from ENT documentation. These findings provide evidence-based guidance for LLM deployment in clinical documentation workflows, suggesting they are best suited for initial extraction with human verification rather than autonomous operation.

14
Benchmarking Open-Source Vision-Language Models for Brain Metastasis Assessment on Single-Slice Contrast-Enhanced MRI

Kim, J.; Kim, B.-s.; Ko, J. S.; Dong, J.; Youn, S. Y.; Jang, J.; Ahn, K.-J.

2026-08-26 radiology and imaging 10.64898/2026.08.24.26361169 medRxiv
Top 0.2%
2.1%
Show abstract

Purpose Open-source vision-language models (VLMs) can be locally deployed without external internet access, potentially enhancing data security. This study compared the diagnostic performance of general-purpose and medical-purpose open-source VLMs and evaluated their ability to characterize brain metastases on contrast-enhanced (CE) MRI. Materials and Methods Sixty lesion-positive axial CE T1-weighted images and sixty matched lesion-negative images from 60 patients were analyzed using three general-purpose VLMs-InternVL3-8B, Qwen2.5-VL-7B-Instruct, and MiniCPM-V-4.5-and three medical-purpose VLMs-MedGemma-4B-it, LLaVA-Med v1.5, and HuatuoGPT-Vision-7B. Lesion detection performance was assessed using sensitivity, specificity, and balanced accuracy. On lesion-positive images, accuracy was evaluated for lesion count, laterality, anatomic location, enhancement pattern, necrosis, vasogenic edema, and mass effect. Model differences were assessed using Cochran's Q tests followed by pairwise McNemar tests with Benjamini-Hochberg correction. Results The median age of the study patients was 67 years (IQR, 61.0-70.5 years), and 35 patients were male (58.3%). MiniCPM-V-4.5 showed the most balanced diagnostic performance, with a sensitivity of 78.3% (95% CI, 66.4-86.9%) and a specificity of 85.0% (95% CI, 73.9-91.9%), and significantly higher balanced accuracy than all other models. Significant overall differences were observed for lesion count, laterality, location, enhancement pattern, necrosis, and mass effect, but not for vasogenic edema (FDR-adjusted P = 0.056). HuatuoGPT-Vision-7B and MedGemma-4B-it showed relatively consistent accuracy across multiple image assessment tasks, although their performance remained modest. Conclusion Our study demonstrated substantial heterogeneity in the performance of open-source VLMs in brain metastasis evaluation, and medical-purpose VLMs did not outperform general-purpose VLMs.

15
NeuroMesh: A Bottleneck Topology Controller for Missing-Modality Brain Tumor Segmentation - A Mechanistic Pilot Study on BraTS

Kamalakannan, N. K.; Kamalakannan, J.

2026-08-27 bioinformatics 10.64898/2026.08.23.746542 medRxiv
Top 0.2%
1.7%
Show abstract

Deep segmentation networks can degrade sharply when an expected MRI sequence is unavailable at inference. We present NeuroMesh, a bottleneck controller that combines a gated recurrent unit (GRU) with a graphconvolutional edge-activation mask, designed to adapt a U-Net-style segmentation backbone to missing input. We evaluate NeuroMesh in a pilot study using a 30-patient subset of the BraTS 2020 benchmark (22 training, 4 validation, and 4 held-out test patients) under a prespecified frozentest protocol. On the frozen test set, NeuroMesh has higher tumor-core and enhancing-tumor Dice than a plain U-Net in most evaluated missing-modality conditions, but wholetumor Dice falls from 0.596 to 0.108 when FLAIR is missing, compared with 0.604 to 0.545 for the plain U-Net. Direct analysis of the predicted edge-activation mask shows negligible change across modality-availability conditions. A parameter-light static-gating control reproduces the FLAIR failure mode without recurrence, a failure-signal input, or graph-structured machinery. These results do not support the intended interpretation that the trained controller performs input-conditional topology rewiring at the scale of this pilot. Instead, they expose a discrepancy between architectural intent and realized behavior and identify a specific missing-modality failure mode that warrants further investigation. Given the small validation and test sets, the findings are descriptive and do not establish clinical or population-level generalization.

16
Quantitative MRI Preprocessing: Effects of Tissue-Specific Smoothing Approaches on Statistical inference

Jacquemin, A.; Phillips, C.

2026-08-27 neuroscience 10.64898/2026.08.24.746651 medRxiv
Top 0.3%
1.2%
Show abstract

Background: Quantitative MRI (qMRI) provides voxel-wise measurements of tissue properties related to myelin, iron and water content, making it a powerful tool for studying brain aging and microstructural alterations in vivo. However, conventional spatial smoothing can introduce partial-volume effects and blur tissue boundaries, potentially affecting both statistical sensitivity and anatomical specificity. Several tissue-specific smoothing strategies have been proposed to address these limitations, yet their relative impact on voxel-wise statistical analyses remains insufficiently characterized. The present study aims (i) to systematically compare three tissue-specific smoothing strategies: a linear tissue-weighted compensated approach (TWS), a generalized version of nonlinear tissue-masked compensated smoothing approach (gTSPOON), and an intensity-weighted edge-preserving approach based on the Smallest Univalue Segment Assimilating Nucleus smoothing (SUSANs), and (ii) to investigate how smoothing approaches interact with statistical inference frameworks by comparing parametric and non-parametric voxel-wise analyse. Methods: Analyses were performed on a publicly available lifespan qMRI dataset comprising 138 healthy participants (19-75 years) and quantitative maps of MTsat, PD, R1, and R2*. The generalized TSPOON (gTSPOON) method was implemented using tissue-specific masks derived from probabilistic tissue segmentation. All three smoothing approaches (TWS, gTSPOON and SUSANs) were parameterized to achieve comparable nominal spatial smoothing. Age-related effects were investigated separately in GM and WM using voxel-wise general linear models following a previously published framework. Statistical inference was assessed using multiple complementary approaches, including parametric Random Field Theory (RFT), under both stationarity and non-stationarity assumptions, as well as non-parametric permutation-based inference. In addition to conventional thresholded statistical parametric maps, voxel-wise log-likelihood (LL) maps were computed to quantify general linear model (GLM) goodness-of-fit independently of statistical thresholding. Bland-Altman analyses and spatial agreement metrics were subsequently used to compare smoothing strategies. Results: TWS and gTSPOON produced highly similar spatial distributions of age-related effects across all qMRI parameters and tissue classes. However, TWS consistently yielded a larger number of significant voxels and clusters, reflecting slightly higher sensitivity, from slightly wider effective smoothness and reduced RESEL counts. By contrast, SUSANs generated substantially fewer significant voxels and clusters, associated with approximately half the effective smoothness and a markedly larger number of RESELs. Despite these differences in statistical sensitivity, voxel-wise LL analyses revealed distinct anatomical preferences for each smoothing strategy. TWS provided the best model fit predominantly within GM, whereas gTSPOON showed superior performance in homogeneous WM regions. Conversely, SUSANs achieved the highest LL values at GM-WM interfaces, particularly within sulcal and gyral transitions, indicating improved preservation of sharp anatomical gradients. These spatial patterns were consistently observed across MTsat, PD, R1 and R2* maps. Comparisons across stationary and non-stationary RFT assumptions revealed only minor differences, while non-parametric inference produced highly concordant results, indicating that the primary source of variability originated from the smoothing procedure itself rather than the inference framework. Conclusions: Tissue-specific smoothing strategies substantially influence both statistical sensitivity and voxel-wise model fitting in qMRI analyses. While TWS and gTSPOON provide highly consistent results, the edge-preserving SUSANs approach preferentially enhances model fit at tissue boundaries. Importantly, voxel-wise log-likelihood mapping revealed that no smoothing strategy is uniformly optimal throughout the brain; instead, each method exhibits anatomically preferential regions where model fit is maximized. These findings suggest that smoothing should be viewed as a region-dependent optimization problem and highlight voxel-wise LL mapping as a principled framework for selecting or developing adaptive smoothing strategies tailored to specific neuroanatomical structures and biological processes, including age-related brain changes.

17
ClinSeg: Robust Brain Segmentation for Clinically Acquired Pediatric MRI

Levitis, E.; Tregidgo, H. F. J.; Zimmerman, D.; Jung, B.; Karandikar, S.; Gardner, M.; Mattisson, P.; Kafadar, E.; Zapaishchykova, A.; Kann, B. H.; Sotardi, S. T.; Vossough, A.; Huang, H.; Billot, B.; Iglesias Gonzales, J. E.; Alexander, D. C.; Alexander-Bloch, A. F.; Seidlitz, J.

2026-09-02 pediatrics 10.64898/2026.08.28.26361643 medRxiv
Top 0.3%
1.1%
Show abstract

Clinical brain MRIs from pediatric health systems represent a viable resource for modeling early neurodevelopmental trajectories and studying neurodevelopmental risk in real-world populations. However, a limitation to date has been the performance of existing segmentation tools for measuring various brain phenotypes in clinical scans. In particular, many tools underperform in infant scans due to morphological and physical changes such as rapid myelination. Here, we introduce ClinSeg: a robust segmentation approach tailored to early-life clinical MRIs with variable orientation, resolution, and contrast. We leverage existing registration and synthetic data generation tools to construct a training corpus for a 3d U-Net spanning anatomical and contrast diversity, including scans with morphological abnormalities from a pediatric hospital. Validated against manual segmentations, ClinSeg outperforms existing models in infancy while matching them in childhood and adolescence. Finally, ClinSeg enables the construction of reference brain growth trajectories in 11,699 individuals from 0-21 years of age, leading to the detection of more nuanced age-related findings in clinical groups.

18
An Explainable and Comparative Transfer Learning Framework for Brain Tumor Classification from MRI Images

Bethala, S.; Vanshika,

2026-08-10 radiology and imaging 10.64898/2026.08.06.26359900 medRxiv
Top 0.3%
1.1%
Show abstract

Automated detection of brain tumors from Magnetic Resonance Imaging (MRI) can accelerate diagnosis and reduce inter-reader variability, yet many existing studies report only top-line accuracy on small datasets, omit efficiency analysis, and provide no interpretability, limiting their clinical credibility. We present a reproducible, comparative, and explainable transfer- learning framework for binary brain-tumor classification. Our framework (i) standardizes a configurable preprocessing pipeline combining CLAHE contrast enhancement and unsharp-mask sharpening, (ii) evaluates a custom CNN baseline and pretrained backbones under an identical training budget, (iii) reports a full metric suite (accuracy, precision, recall, F1, ROC-AUC, PR-AUC, parameter count, and inference latency), and (iv) applies Grad- CAM for spatial interpretability. On a public 253-image MRI dataset (38-image held-out test set), MobileNetV2 achieves the best overall performance (94.74% accuracy, 0.994 ROC-AUC, 0.996 PR-AUC) with only 2.59M parameters and 5.9 ms per- image inference, making it the most deployment-friendly model. Larger backbones (Xception, EfficientNetB0) and the custom CNN converge to degenerate all-positive predictions under the same limited budget, illustrating the small-data overfitting risk that accuracy-only reporting conceals. Grad-CAM confirms that the best model attends to the tumor region. All source code, con- figuration files, and trained evaluation scripts are publicly avail- able at https://github.com/blck-iris/explainable-brain-tumor-mr

19
Perceived Hearing Symptoms Organize the Ear-Disease Comorbidity Network but Are Not a Causal Lever for Brain Health: A Triangulated Analysis

Chen, M.; Huang, Y.; Yu, R.; Xie, Y.; Chen, F.; Huang, J.; Zhao, J.; Ma, Z.; Ma, Z.; Jiang, L.

2026-08-14 otolaryngology 10.64898/2026.08.13.26360183 medRxiv
Top 0.4%
1.1%
Show abstract

Background: Hearing loss is a potentially modifiable risk factor for brain health, but whether it acts as a causal lever remains unclear. Methods: We constructed an ear-disease comorbidity network from NHANES 2011-2020 (N=18,939, 16 nodes, 62 edges), performed bidirectional Mendelian randomization (MR) across 24 exposure-outcome pairs, and triangulated evidence with longitudinal data from CHARLS (N=17,101). Results: Subjective hearing symptoms (prevalence 6.1%) occupied hub positions in the comorbidity network, whereas objective hearing impairment (8.0%) was sparsely connected. All forward MR estimates were null after multiple-testing correction (IVW P>.05 for 9 of 9 pairs). Reverse MR showed one nominally significant association (cognition to objective hearing beta=-0.15, P=.013) that did not survive correction. Longitudinal analysis yielded HR=1.57 (P=.00004) for subjective hearing symptoms predicting incident depression. Conclusions: Perceived hearing symptoms organize the ear-disease comorbidity network but are not a causal lever for brain health. These findings support a "flag, not lever" framework: subjective hearing symptoms warrant clinical attention as markers of systemic multimorbidity rather than intervention targets for dementia prevention.

20
Modular Arrays for High Precision Wearable MEG

Alexander, N. A.; Mariola, A.; Puvvada, S.; Bezsudnova, Y.; Tierney, T. M.; Barnes, G. R.; Callaghan, M. F.

2026-08-24 neuroscience 10.64898/2026.08.19.745485 medRxiv
Top 0.4%
1.0%
Show abstract

Optically pumped magnetometers (OPMs) can be used for magnetoencephalography (MEG) with equivalent or improved signal to noise ratio, relative to cryogenic MEG, when sensors are placed close to the scalp. OPM-based MEG can also be used in mobile contexts if sensors are placed in lightweight, wearable arrays. Individually tailored, rigid helmets known as scannercasts are currently the only method capable of achieving on-scalp, mobile recordings with high precision. However, these scannercasts are expensive to produce, require structural imaging in advance of the experiment, and can incur lengthy downtime while sensors are transferred between scannercasts. Here, we introduce a solution to these challenges that retains the advantages of scannercasts. We provide detailed steps for constructing a modular, cap-based design, suitable for all head sizes. Using simulations, we compare the leadfield power of this array against an idealised array and a commercially available mobile solution. We then validate our proposed solution empirically, in five participants, and provide a complete data preparation and analysis pipeline. Our design expands the accessibility of OPM-based MEG, and increases participant throughput to levels comparable to other imaging modalities. Crucially, it removes the trade-off between signal quality, mobility and practicality, promoting the unique potential of OPM-based MEG as a tool for studying naturalistic behaviour, and clinical assessment with high precision.